Papers with evaluation error

2 papers
Poller: Are LLMs Suitable for Evaluating Poetry Understanding Task? (2026.findings-acl)

Copied to clipboard

Challenge: Traditional methods for poetry evaluation are expensive and unsuitable for large-scale data.
Approach: They propose a method leveraging Large Language Models to evaluate poetry understanding tasks using Large Language models.
Outcome: The proposed method reduces the evaluation error between LLMs and humans by adopting the poet's perspective.
SciCompanion: Graph-Grounded Reasoning for Structured Evaluation of Scientific Arguments (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing retrieval-augmented generation (RAG) methods fail to provide deep, relational understanding of scientific literature.
Approach: They propose a graph-grounded reasoning framework for structured scientific evaluation that uses multi-hop reasoning to iteratively construct contextual graphs and generate structured critiques.
Outcome: The proposed framework reduces evaluation error by over 30% compared to baselines and allows smaller models to outperform larger models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations